Back

European Journal of Human Genetics

Springer Science and Business Media LLC

Preprints posted in the last 30 days, ranked by how well they match European Journal of Human Genetics's content profile, based on 58 papers previously published here. The average preprint has a 0.04% match score for this journal, so anything above that is already an above-average fit.

1
A Method to Analyze Low-Quality Archaic Human Genomes and its Application to the Teshik-Tash 1 Neandertal

Sümer, A. P.; Iasi, L. N. M.; Bossoms Mesa, A.; Slon, V.; Essel, E.; Hajdinjak, M.; Zorn, J.; Schmidt, A.; Nagel, S.; Nickel, B.; Viola, B.; Ziganshin, R.; Buzhilova, A.; Derevianko, A.; Pääbo, S.; Peter, B. M.

2026-08-14 genetics 10.64898/2026.08.10.743885 medRxiv
Top 0.1%
12.0%
Show abstract

The Teshik-Tash 1 child whose remains were found in Uzbekistan represents the southeastern-most extent of the known Neandertal range, providing an important link with the better studied Caucasus and Altai Mountain ranges. However, due to poor DNA preservation, studying the genetics of Teshik Tash 1 has remained elusive. Here we present analyses of the nuclear DNA from the Teshik-Tash 1, from extracts that are highly contaminated with present-day human DNA. To achieve this, we developed a new computational method, admixslug, that jointly models contamination and population relationships, in order to infer the relationship of a target individual from which only low-quality nuclear DNA is available, to high-quality archaic human genomes. After validating admixslug, we show that Teshik-Tash 1 is genetically more similar to later Neandertals from Western Eurasia than to older Neandertals from the Altai Mountains. We estimate that Teshik-Tash 1 split from the Western Eurasian lineage between 80,000 and 100,000 years ago. Despite the geographical proximity of Teshik-Tash 1 to the Denisovan range, we find no evidence for Denisovan ancestry in his genome. Our results demonstrate that admixslug enables the study of archaic human specimens in cases where DNA preservation was previously considered too poor for population genetic analyses.

2
Readability Assessment of Patient-Reported Measures Used During Heritable Cancer Genetic Testing

Adegbesan, A. C.; FitzGerald, L.; Dickinson, J. L.; Raspin, K.; Roydhouse, J.

2026-08-17 oncology 10.64898/2026.08.13.26360322 medRxiv
Top 0.1%
8.2%
Show abstract

Background: Patient-reported measures (PRMs), including patient-reported outcome and experience measures, capture patients perspectives on their health status and healthcare experiences. In cancer genetics, PRMs have been used to assess genetic knowledge, psychosocial outcomes, and decision-making. However, patients must understand these measures to provide useful information, an ability which is influenced by general and health literacy levels. Readability guidelines recommend that patient-facing materials be written at or below a Grade 6 level. This study evaluated the readability of PRMs used in a cancer genetic testing context. Objective: To assess whether PRMs used in heritable cancer genetic testing meet recommended readability levels using validated indices. Methods: PRMs were identified from a recent systematic review of PRMs used in heritable cancer genetic testing, which reported 83 instruments across eight categories. English-language PRMs containing structured question items and response scales were eligible for extraction and converted into plain text for analysis. Readability was assessed using four validated indices: Flesch Kincaid Grading Level (FKGL), FORd, CAylor, and STicht (FORCAST) formula, Flesch Reading Ease Score (FRES), and Simple Measure of Gobbledygook (SMOG) via an automated readability software. Descriptive analysis and numerical comparison evaluated readability levels across PRM categories and against the recommended Grade 6 reading level. Results: Sixty-five PRMs met the eligibility criteria, with most, including validated instruments, exceeding the recommended Grade 6 reading level. Across the eight categories, genetics-specific PRMs required the highest readability levels, indicating higher readability demands. Conclusions: Most PRMs, particularly those specific to genetics, do not meet readability guidelines. This may limit their accessibility to individuals with limited general and health literacy. Development of PRMs specific to genetics should consider strategies to improve readability, such as plain-language approaches and involvement of individuals with limited general or health literacy. Keywords: readability, patient-reported measures, cancer, genetic testing, health literacy

3
A Randomized Non-Inferiority Trial of an eHealth Delivery Alternative for Cancer Genetic Testing for Hereditary Cancer (eREACH2)

Lee, K. T.; Egleston, B.; Fetzer, D.; Domchek, S. M.; Fleisher, L.; Wen, K.-Y.; Wagner, L.; Roberts, S.; Howe, S.; Cacioppo, C.; Christiansen, J.; Karpink, K.; Selmani, E.; Mastaglio, E.; Weinberg, M.; Wood, E. M.; Feng, J.; John, S.; Schweickert, K.; Mcleod, B.; Bradbury, A. R.

2026-09-03 genetic and genomic medicine 10.64898/2026.09.01.26361920 medRxiv
Top 0.1%
6.9%
Show abstract

Background: Many at-risk patients lack access to genetic services due to a genetic counselor (GC) workforce shortage. Little is known about how digital alternatives impact patients with and without cancer who meet criteria for genetic testing. Methods: eREACH2 is a randomized 4-arm non-inferiority trial where pre-test (visit 1) and/or return of results (visit 2) GC counseling was replaced with a patient-centered digital intervention. Arms include: A (GC/GC), B (GC/digital), C (digital/GC) and D (digital/digital). Primary outcomes were non-inferiority in uptake of genetic services and change in genetic knowledge and general anxiety from baseline to post-disclosure of results (T0-T2). Secondary cognitive and affective outcomes were assessed using non-inferiority ANOVAs and equivalency chi-squared tests in intention-to-treat and per-protocol analyses. Findings: 773 participants were recruited nationwide; 46.6% from rural areas. Mean age was 51 years (range 20-87), 13% male, 12% non-white, 29% had less than a college education, and 33% had a personal history of cancer. 584 (76%) patients completed testing (14% had a positive result, 16% had a VUS). In the primary ITT analyses, we met the non-inferiority for uptake of genetic services and anxiety, but results were inconclusive for knowledge. Secondary outcomes were heterogeneous across arms. Arm C demonstrated consistently favorable effects, while Arms B and D showed less favorable outcomes in select domains (e.g. satisfaction and MICRA). Patients who received positive or VUS results via digital disclosure had significantly higher MICRA scores - indicating greater negative response to testing. Interpretation: In this large, randomized trial of patients with and without cancer, the eREACH intervention was effective for pre-test counseling, but inconclusive for digital disclosure of results. Exploratory analyses suggest that digital delivery could be a reasonable alternative for individuals receiving negative results, while those receiving positive or VUS results may derive some short-term psychosocial benefit from GC disclosure.

4
Patient and Clinician Perspectives on Centralized Cascade Screening for Familial Hypercholesterolemia in the United States: A Qualitative Implementation Study

Roberts, M. C.; Jones, L. K.; Brown, A.; Carda-Auten, J.; Cuchel, M.; Hilton, A. R.; Khera, A.; Rothstein, M.; Soe, K.; Sullivan, A.; Tricou, E.; Vu, M. B.; Weintraub, W. S.; Ahmad, Z.

2026-08-19 cardiovascular medicine 10.64898/2026.08.18.26359725 medRxiv
Top 0.1%
5.7%
Show abstract

Objective: To identify patient- and clinician-reported barriers, facilitators, and design requirements for a centralized cascade-screening program for familial hypercholesterolemia (FH) in the United States. Methods: From June through November 2023, we conducted individual telephone interviews with 20 patients with FH and 10 clinicians recruited from UT Southwestern Medical Center, Parkland Health, the North Texas Veterans Affairs, and other clinical settings. Interview guides were informed by the Consolidated Framework for Implementation Research. Transcripts were coded in Dedoose using a piloted codebook, with discrepancies and emergent themes resolved through consensus. An advisory panel then helped translate interview findings into program design requirements and implementation strategies. Results: Five themes characterized barriers and facilitators to centralized cascade screening: (1) health-system access and fragmentation, including screening and treatment costs, transportation, and cross-system coordination; (2) privacy and trust, including concerns about genetic information and unsolicited outreach; (3) family relationships and practical burden, including competing demands, language barriers, limited contact, fear, and denial; (4) clinician capacity and workflow, including limited time, knowledge, and genetic-counseling capacity; and (5) communication and care continuity. Participants recommended proband pre-notification of relatives, culturally and linguistically responsive materials, secure data exchange, standardized scripts, flexible testing pathways, and centralized coordination. These findings informed a program model incorporating a secure pedigree platform, educational and communication resources, testing coordination, and linkage to follow-up care. Conclusions: Patients and clinicians identified multilevel determinants that a centralized FH cascade-screening program must address. The findings support specific design requirements but do not establish program feasibility or effectiveness, which require prospective evaluation.

5
Expanding reproductive genetic screening through the inclusion of perinatal treatability

Tan, T. Y.; Haas, S.; Gao, X.; Li, J.; Araji, S.; Liu, A.; Wimberly, C.; Gold, N.; Rentas, S.; Duyzend, M.; Walsh, K. M.; Cohen, J. L.

2026-08-27 genetic and genomic medicine 10.64898/2026.08.24.26361139 medRxiv
Top 0.2%
4.4%
Show abstract

Various professional organizations recommend screening prospective parents for autosomal recessive (AR) and X-linked (XL) conditions, which is reflected in commercial screening panels. There is merit to developing a distinct reproductive gene-list and analytic framework inclusive of genes based on available perinatal intervention, defined as possible prenatal intervention (including investigational) for the fetus or necessary early initiation of approved postnatal treatments. We evaluated a reproductive genetic screening framework that incorporates perinatal actionability across AR, XL, and selected autosomal dominant (AD) genes. Using a curated list of genetic conditions with perinatal intervention, we evaluated five subset gene lists to determine the individual-level number-needed-to-screen (NNS) to identify one individual with at least one qualifying heterozygous variant, defined as a heterozygous pathogenic or likely pathogenic (P/LP) variant in a gene on the specified list. To conduct NNS analyses, we sourced carrier frequency and allele frequency data for each gene and their respective ClinVar-curated high-confidence (>=2 star) P/LP variants, from two population databases -- gnomAD v4.1 and All of Us (AoU) v8. The analyses produced an individual-level NNS of 3.20 (CI: 3.193, 3.212) using gnomAD and 3.62 (CI: 3.606, 3.640) using AoU. These estimates do not represent couple-level reproductive risk, affected-pregnancy yield, clinical diagnostic yield, or validation of a clinical screening test. These findings support further evaluation of a perinatal-actionability framework, with clinical value dependent on which genes drive yield, and whether the relevant gene, variant, mechanism, and phenotype combinations are actionable in a reproductive or perinatal context for both the pregnant woman and her future offspring.

6
A Framework For Large-Scale Reconstruction Of Extended Pedigrees To Facilitate Gene Discovery In ALS

van Oosten, D.; Beele, P.; Wang, B.-n.; Plasmans, S. J.; Wolthuis, N.; van den Berg, K.; Blom, M. P. T.; Meyjes, M.; van der Schoot, N. D.; Vergunst-Bosch, H.; Kok, A. R.; van der Ven, L. J.; van Es, M. A.; van den Berg, L. H.; Veldink, J. H.; van Rheenen, W.

2026-08-27 genetic and genomic medicine 10.64898/2026.08.21.26360249 medRxiv
Top 0.3%
3.9%
Show abstract

Importance: With emerging gene-targeted therapies in amyotrophic lateral sclerosis (ALS), gene discoveries and genetic diagnoses provide a crucial path to treatment. Pathogenic variants with moderate effect or incomplete penetrance, however, remain unidentified in genome-wide association studies and can appear sporadic in small modern-day pedigrees. Lack of recognition of familial clustering of ALS, in turn, limits opportunities for gene discovery, genetic diagnosis, risk counseling, and treatment. Objective: To determine the power of automated reconstruction of extended pedigrees, integrating archive records and genetic relatedness, in gene-discovery studies. Design: Retrospective observational study of Dutch ALS patients with the C9orf72 hexanucleotide repeat expansion (HRE), combining clinical family history, civil records, and genome-wide genotyping for relatedness and identity-by-descent (IBD) inference. Setting: National, population-based ALS cohort from the Netherlands and digitized population archives enabling systematic reconstruction of extended pedigrees. Participants: Individuals with ALS and a confirmed C9orf72 HRE. Participants must have provided a clinical family history and traceable Dutch ancestry documented in population archives. Main Outcomes and Measures: The primary outcome was the proportion of C9orf72 HRE carriers with newly identified (distant) relatives with ALS compared with clinical family history. The secondary outcome was the precision of IBD-based methods to fine-map the C9orf72 HRE. Other outcomes included phenotypic similarities between distantly related patients. Results: Among 238 C9orf72 HRE carriers, 91 could be included in one of 39 extended pedigrees dating back to ~1800, with relationships up to the eighth degree of relatedness. Compared with clinical family history alone, our approach increased the number of identified relationships by 2.5-fold. Genome-wide IBD analysis revealed shared haplotypes encompassing the C9orf72 HRE in 94% of pedigrees by [≥]7 meioses in 25.7-127.8 centimorgans total IBD shared. Conclusions and Relevance: Large-scale interrogation of archives facilitates reconstruction of extended pedigrees for ALS patients carrying the C9orf72 HRE. This combined genealogical-genetic approach supports the reclassification of apparently sporadic cases, facilitates the discovery of new disease-causing variants in ALS, and is generalizable to other late-onset neurodegenerative diseases. Automated pedigree reconstruction from genealogical data and visualization in an interactive databrowser are implemented in the open-source Mangrove software.

7
Evaluating the Impact of Principal Component and Mixed Model Approaches on Polygenic Risk Score Portability to Diverse Ancestries in the UK Biobank

Harikrishnan, A. S.; Kelly, C. M.

2026-08-19 genetic and genomic medicine 10.64898/2026.08.17.26360388 medRxiv
Top 0.3%
3.2%
Show abstract

Polygenic risk scores (PRS) offer considerable potential for precision medicine. How ever, their predictive performance often attenuates when applied to populations that differ from the genome-wide association study (GWAS) training population. There are many potential sources of this portability problem, and one relatively under-explored contributor is the presence of residual confounding in GWAS summary statistics. In particular, confounding specific to the training population may contribute to predictive performance that does not transfer to other populations, such that improved control of population stratification could potentially improve PRS portability. Here, we investigated whether varying levels of population stratification adjustment, through the inclusion of principal components and the use of mixed models, altered PRS portability in three broad ancestry groups in the UK Biobank. The PRS were built using European training data for coronary artery disease and type 2 diabetes and subsequently evaluated in South Asian, African, and Latin American participants. We found that increasing PC adjustment did not produce a consistent trend in portability across ancestry groups or phenotypes, despite modest reductions in the LDSC intercept. However, substantial ancestry- and phenotype-specific effects on transferability were observed. Mixed-model association provided no significant change in PRS discrimination or portability. These findings highlight the need for a better understanding of the nature of residual confounding in PRS and whether improving the causal validity of GWAS results can ultimately improve the transferability of predictive accuracy between populations.

8
Genetic Architecture and Sample Size Impact Relative Performance of Nonlinear Machine Learning and Standard Polygenic Risk Scores

Zhu, J.; Baousi, A.; Morris, A. P.; Guo, H.

2026-09-03 genetic and genomic medicine 10.64898/2026.08.29.26361109 medRxiv
Top 0.3%
3.2%
Show abstract

Standard polygenic risk scores (PRSs) are constructed based on additive genome-wide association study (GWAS) summary statistics. Nonlinear machine learning methods have been increasingly applied to construct PRSs directly from individual-level data, with the aim of improving predictive performance over standard PRSs through their ability to model non-additive genetic effects. However, their superiority across studies has been inconsistent, and the conditions under which they provide meaningful improvements remain unclear. We combined theoretical analysis, simulations and a real-world application to investigate when two widely used nonlinear machine learning methods, random forest and XGBoost, outperform standard PRSs. Theoretical analysis showed that standard PRSs can implicitly capture part of the genetic variance attributable to nonadditive genetic effects through their contributions to marginal SNP effects, thereby losing less information than commonly assumed. Although nonlinear models have a higher theoretical potential, their greater flexibility incurs a bias-variance trade-off that can limit predictive gains at finite sample sizes. Simulations showed that XGBoost outperformed the standard PRS only when the genetic architecture involves a sufficiently large proportion of interaction genetic variance concentrated across relatively few interaction effects and large training samples were available. Random forest consistently underperformed the standard PRS. In an application to ischemic heart disease prediction using UK Biobank data, XGBoost showed no meaningful improvement in predictive performance over the standard PRS, whereas random forest again performed worse. Together, these findings suggest that nonlinear machine learning do not uniformly outperform standard PRSs; rather, their relative performance depends jointly on genetic architecture and training sample size. Our study helps to reconcile the inconsistent results reported across previous studies and provides a framework for identifying settings in which more complex PRS models are likely to be beneficial.

9
A meta-analysis of ancient and present-day Central Eurasian genome data to revise archaic hominin ancestry

Rymbekova, A.; Kuhlwilm, M.

2026-08-14 genomics 10.64898/2026.08.10.743976 medRxiv
Top 0.3%
3.1%
Show abstract

Archaic introgression has shaped the evolutionary history of Eurasian populations, yet Central Eurasian region remains understudied despite being at the crossroads of ancient human migration. Here, we analyzed the whole-genome data of five Central Eurasian (CE) individuals from Early Bronze Age (EBA) and five present-day CE individuals to characterize the archaic introgression landscape. We estimated that archaic introgression from Neanderthal and Denisovan archaic hominins comprises approximately 2.2% of the Central Eurasian genomes. Both amount and chromosomal distribution of archaic introgression remained largely unchanged between the EBA and present-day CE individuals. Putative introgressed fragments matching the Altai Neanderthal and the Altai Denisovan were retrieved. Our results suggest that while the archaic introgression levels seemingly remained stable over the past several thousand years, larger modern CE genomes panels will be required to fully characterize the genomic landscape of archaic ancestry in the region.

10
Benchmarking Twist Genotyping-by-Sequencing Against Whole-Genome Sequencing in Nuclear Families

Klugerman, J.; Iossifov, I.; Ye, K.

2026-08-06 bioinformatics 10.64898/2026.07.31.742127 medRxiv
Top 0.4%
2.4%
Show abstract

Genome-wide genotyping is widely used in human genetics research, and targeted sequencing-based approaches such as the Twist Bioscience genome-wide SNP capture platform (GxS) have emerged as alternatives to conventional SNP arrays. Here, we evaluated GxS genotype calls from 555 individuals in 184 nuclear families against matched whole-genome sequencing (WGS) calls and compared platform performance with that of the Illumina Infinium Global Screening Array-24 (GSA), which was evaluated in 987 individuals from 279 nuclear families. Genotype data were harmonized across platforms, and analyses were restricted to overlapping SNP loci. Across all callable positions, mean per-SNP call rates were 98.26% for GxS and 98.67% for GSA. Overall SNP concordance with WGS was 99.79% for GxS and 99.87% for GSA, and mean per-individual concordance was also 99.79% and 99.87%, respectively. Per-trio Mendelian violation rates of GxS are about 10 times those of WGS, while those of GSA are about 4 times those of WGS on average. These results indicate that GxS performs slightly worse than GSA by key concordance and inheritance metrics, while still showing strong overall agreement with WGS.

11
Using summary data to detect and quantify ascertainment in biobanks

Olasege, B. S.; Campos, A. I.; Sidorenko, J.; Lin, T.; Barry, C.-J. S.; Maseras, G. T.; Vilhjalmsson, B. J.; Wray, N. R.; Hivert, V.; Yengo, L.

2026-08-07 genetics 10.64898/2026.08.02.742371 medRxiv
Top 0.5%
2.3%
Show abstract

Non-random participation in genetic studies can bias associations between genetic variants and outcomes. Existing methods to detect ascertainment bias often require individual-level data, thus limiting their broad applicability. Here, we introduce a summary-statistics-based method to detect and quantify ascertainment bias in large-scale genetic studies. Our method estimates a parameter,{theta} , which captures deviations in the mean polygenic score (PGS) of an ascertained sample relative to its expectation across non-ascertained or differentially ascertained references. We show through extensive simulations that our method is robust to population stratification and reference misspecification unlike naive mean PGS comparison. When applied to 21 traits across 11 large-scale biobanks, our method recapitulates known patterns of ascertainment and detects new evidence of ascertainment on genetic susceptibility to depression, height and blood pressure in many biobanks. Overall, our framework enables systematic assessment of ascertainment directly from summary statistics and provides a scalable tool for evaluating representativeness in large scale genetic studies.

12
PANACEA: a framework to maximise genetic diversity in genome-wide association study meta-analyses

Yap, C. F.; Morris, A.

2026-08-10 genetic and genomic medicine 10.64898/2026.08.06.26359891 medRxiv
Top 0.5%
2.2%
Show abstract

There have been recent efforts by the human genetics research community to increase the genetic diversity of participants contributing to genome-wide association studies (GWAS) of complex human traits and diseases. The traditional multi-ancestry GWAS approach is to first assign participants to continental ancestry labels based on their genetic similarity to individuals in reference datasets. Ancestry-specific GWAS are then conducted separately for each continental label, the results of which are aggregated through multi-ancestry meta-analysis. However, with this approach, a participant may be assigned to an ancestry group that does not reflect their personal view of ethnicity/race or may be excluded because their genetic ancestry is not sufficiently similar to individuals in reference datasets to be assigned to a single group. Here, we present a novel pipeline (PANACEA) for fully inclusive multi-ancestry meta-analysis that employs a continuous and multi-dimensional representation of ancestry that maximises the genetic diversity of GWAS. Through application to multi-ancestry GWAS of type 2 diabetes susceptibility and simulations, we demonstrate that the inclusive pooled analysis provides equivalent protection against population structure to a traditional ancestry-stratified analysis but, importantly, offers increased power to detect association through increased sample size by not excluding participants with outlying ancestry. The pooled inclusive analysis also enables assessment of ancestry-correlated heterogeneity in allelic effects without the need to assign participants to continental labels that may not sufficiently reflect genetic diversity within ancestry groups.

13
Integrative optical genome mapping and long-read sequencing resolve constitutional complex rearrangements at nucleotide resolution

Burssed, B.; van der Sanden, B.; Hops, W.; Neveling, K.; Kamping, E.; van Beek, R.; den Ouden, A.; Derks, R.; Timmermans, R.; Perrone, E.; Ramos, M. A.; Bellucco, F. T.; Hoischen, A.; Melaragno, M. I.

2026-08-28 genomics 10.64898/2026.08.27.747510 medRxiv
Top 0.5%
2.0%
Show abstract

Complex rearrangements are one of the rarest types of structural variants (SVs) and can be divided into two categories: complex chromosomal rearrangements (CCRs) and complex genomic rearrangements (CGRs). CCRs include structural rearrangements that present at least three breakpoints and show exchange of genetic material between more than two chromosomes and CGRs are rearrangements that present more than one junction and/or more than one SV in cis. They are usually formed by one of the chromoanagenesis mechanisms, where a massive disruptive cellular event leads to multiple structural rearrangements. Classical cytogenomic techniques have been commonly applied for their characterization, but methodologies that involve longer DNA molecules, namely optical genome mapping (OGM) and long-read genome sequencing (lrGS), present a considerably higher SV detection resolution, revealing more details about the rearrangements, including precise breakpoint location. Here, we describe six patients with complex rearrangements investigated through a combination of different techniques: karyotyping, chromosomal microarray, and OGM were performed to characterize the rearrangements. Subsequently, lrGS was used to further resolve the alterations, refine their breakpoints' location, and sequence their junction points. Three patients presented CCRs involving three, four, and six chromosomes, while three exhibited CGRs involving one different chromosome each, providing a variety of complex SVs to show the importance of each technique and their combination in rearrangement resolution. In total, the complex rearrangements presented 127 breakpoints, 66 junction points and involved 14 of the 24 chromosomes. Higher-resolution techniques revealed additional complexity in all cases. Despite the advances provided by OGM and lrGS, conventional karyotyping remained indispensable for complete rearrangement resolution. In two patients, the findings supported a novel mechanism combining features of the different chromoanagenesis processes. Furthermore, evidence of inherited alterations was identified, and the comprehensive characterization of the rearrangements enabled more accurate genotype-phenotype correlations. Our findings indicate that an integrated approach combining karyotyping, OGM, and lrGS can completely resolve SVs, including complex rearrangements.

14
Robertsonian translocations in Danish sika deer (Cervus nippon). Markers for absent F1-hybridization with red deer (C. elaphus) and implications for selection, speciation and infertility

Tommerup, N.; Alsing, K. K.; Budtz-Jorgensen, E.; Thune-Stephensen, F.; Ingstrup, A. J.

2026-08-18 genetics 10.64898/2026.08.10.743091 medRxiv
Top 0.5%
1.9%
Show abstract

EU has reclassified the sika deer (Cervus nippon) as an undesirable invasive species based on reports that hybridization with the indigenous red deer (C. elaphus) may produce fertile offspring. Since sika-derived DNA previosuly introduced into the red deer population (introgression) cannot be removed, the crucial question is whether new (F1) hybridisation occur. To address this, we analysed the chromosomes in 56 sika and 22 red deer. All red deer had a chromosome number 2n=68. In contrast, the chromosome number in sika ranged from 64 to 67, due to the variable presence of two sika-specific Robertsonian translocations (ROB1,ROB2). In the free-ranging sika population in Jutland, >90% of the sika deer were homozygote for at least one of these ROBs, excluding that they could be F1-hybrids. Moreover, ROB2 was in Hardy-Weinberg equilibrium, further supporting the absence of gene flow between the two species. In contrast, ROB1 was in Hardy-Weinberg disequilibrium, suggesting negative fitness of heterozygotes, including potential F1-hybrids. In Jaegersborg Deer Park, the eight examined sika deer had the same genotype (absence of ROB1, homozygosity of ROB2), supporting that it is a founder population which may have been isolated for [~]100 years. Again, none of these can be F1-hybrids due to the homozygosity of ROB2. We conclude that F1-hybridisation between sika and red deer either does not occur or occur very rarely in Denmark. The study establish the Danish sika-populations as unique models for adressing important biological questions: What underlies the absence of hybridisation? Why are ROBs frequent in sika deer but not in the closely related red deer? How fast do new species/subspecies develop in isolated founder populations? Which factors determine, that some ROBs have little heterozygous effects, whereas others are selected against, with implications for the role of ROBs as genetic barriers promoting speciation, and for fertility problems in some human ROB carriers.

15
Sparse sampling and rare-variant depletion distort PCA visualizations of population structure: recovery with objective-guided manifold learning

Koci, J.; Flegontova, O.; Changmai, P.; Vyazov, L. A.; Cooper, L. R.; Ashrafikarahroudi, S.; Sencan, Z.; Flegontov, P.

2026-08-13 genetics 10.64898/2026.08.11.744230 medRxiv
Top 0.5%
1.9%
Show abstract

Principal component analysis (PCA) is routinely used to visualize population structure, yet how sparse sampling and rare-variant depletion affect low-dimensional plots remains poorly understood. Using spatial simulations, we show that these factors interact to distort visualization of genetic landscapes, producing triangular and three-ray patterns, artificial outliers and misleading clines. We develop an objective-guided manifold-learning framework that searches across genotype normalization, PCA representation and dimensionality, distance metrics, and UMAP, densMAP and PHATE parameters. High-dimensional classic PC scores consistently outperform the eigenvectors used in population genetics, but other optimal parameters and ranking objectives depend on data quality, sampling and SNP ascertainment. Across six human and animal datasets, optimized embeddings recover fine-scale structure obscured by PCA and supported by independent genetic evidence. In ancient Eurasia, optimized PHATE resolves Slavic-associated structure corroborated by haplotype-sharing communities, qpAdm, and Y-chromosome lineages. These results call for caution in interpreting PCA plots and establish optimized manifold learning as a hypothesis-generating approach.

16
Locus-specific gene-context interactions improve polygenic prediction

Fonseca, R.; Caggiano, C.; Costantino, M.; Dominguez, O.; Kenny, E.; Dahl, A.

2026-08-28 genetics 10.64898/2026.08.26.746823 medRxiv
Top 0.6%
1.9%
Show abstract

Polygenic scores (PGS) are a primary output of large-scale genetic studies and are being deployed in clinical and non-clinical settings. However, current PGS assume simple additive models that ignore context-specific genetic effects, which likely reduce their accuracy and robustness. To address this, we developed PGSC, a PGS framework to incorporate locus-specific gene-context interaction effects (GxC). Simulations show PGSC is robust under the additive model and outperforms PGS in realistic settings. Using sex, age, and statin treatment status as contexts in UK Biobank, we find that PGSC outperforms PGS on average across 48 traits, with substantial improvement in some cases, such as GxSex for testosterone, GxAge for bilirubin, and GxStatins for LDL cholesterol. PGSC consistently outperforms a simple genome-wide GxC model, ampPGS, which only outperforms PGS when a context uniformly amplifies all genome-wide additive effects. Critically, PGSC improvements replicate across ancestries in the UK Biobank and in an external cohort, the Mount Sinai Million Health Discovery Program. Finally, we test robustness to log-scale phenotypes and find that ampPGS gains vanish, while the locus-specific GxC components in PGSC persist. Overall, PGSC is a simple, robust framework that demonstrates GxC effects can improve out-of-sample PGS prediction and is a step toward precision treatment.

17
Reduced PDE4D expression and activity in Acrodysostosis Type 2 patient fibroblasts underlie disease pathology

Gardner, O. F.; Ling, J.; Munkongcharoen, T.; Kyurkchieva, E.; Leitch, H. G.; Wilson, L. C.; Baillie, G. S.; Ferretti, P.

2026-08-11 cell biology 10.64898/2026.08.10.743905 medRxiv
Top 0.6%
1.8%
Show abstract

BackgroundAcrodysostosis type 2 (ACRDYS2) is a rare autosomal dominant disease characterized by skeletal defects and cognitive deficit, with clinical symptoms observed in multiple other tissues including the skin. It is caused by mutations in a phosphodiesterase, PDE4D, a key regulator of cAMP/PKA (cyclic adenosine monophosphate / protein kinase A) signalling. Despite its well-defined genetic causes, the molecular mechanisms underlying the disease remain poorly understood, with studies based largely on engineered cellular models reaching conflicting interpretations. MethodsTo investigate how endogenous dynamics are affected by PDE4D mutations in unmanipulated cells, we studied PDE4D transcript and protein expression, activity and downstream signalling in native dermal fibroblast from ACRDYS2 patients and healthy controls. ResultsSignificant reduction in total PDE4D expression in patient cells was observed both at the transcript and protein level, with marked decreases in the long isoforms PDE4D4 and PDE4D7; a reduction in PDE4D9 mRNA was also observed. PDE4D enzymatic activity was reduced in ACRDYS2 fibroblasts, though total PDE activity was largely preserved. Reduced PDE4D expression was associated with an increase in the phosphorylated form of the cAMP-responsive transcription factor CREB and elevated PRKAR1A (PKA type 1 regulatory subunit alpha) transcript levels, suggesting altered downstream signalling. Interestingly, expression of the related phosphodiesterase family member PDE4B was increased, consistent with a compensatory response to reduced PDE4D function. ConclusionsThis is the first study demonstrating reduced PDE4D expression and isoform-specific dysregulation in native ACRDYS2 cells. Together, our results support a model in which reduction in PDE4D activity and compensatory changes in other PDE4 family members contribute to the molecular pathology of ACRDYS2, providing new insights into the molecular mechanisms underlying this disorder.

18
Uncovering High-Order Epistatic Interactions in GWAS via a Machine Learning-Based Feature Engineering Framework

Byun, J.; Saha, D.; Han, Y.; Shaw, V. R.; Siminovitch, K.; Amos, C. I.

2026-08-09 genomics 10.64898/2026.08.03.742638 medRxiv
Top 0.6%
1.7%
Show abstract

BackgroundGenome-wide association studies (GWAS) often fail to identify higher-order epistatic interactions that contribute to complex inheritance patterns of traits and diseases. While machine learning (ML) can capture non-linear relationships, extracting interpretable insights from these models remains a challenge. We propose a novel tree-based feature engineering framework that uses Classification and Regression Trees (CART) to explicitly encode high-order interaction decision paths as dummy variables. We investigate three path-based encoding strategies: (i) all decision paths, (ii) leaf-node paths only, and (iii) internal-node paths only. This approach aims to transform complex decision boundaries into discrete features that capture nonlinear interactions that are not readily captured by traditional association models. ResultsThe framework was evaluated using genetic data for ANCA-associated vasculitis (AAV). To manage the high dimensionality of the engineered feature space, we applied a comprehensive suite of ML methods across three tasks: (1) Ensemble Learning (Random Forest, XGBoost, and Gradient Boosting Machine); (2) Decision Tree Analysis (CART); and (3) Regression and Classification Tasks (Regularized Linear Regression/LASSO, Support Vector Machine, and Logistic Regression). Stepwise feature selection and regularization were employed to isolate the most informative interaction patterns. Results indicate that incorporating CART-derived interaction paths--particularly those from high-impact regions of the tree--significantly improves classification accuracy and model interpretability compared to using the original feature space alone. ConclusionsThe proposed framework provides a robust, scalable methodology for identifying high-order genetic interactions. By bridging the gap between the predictive power of ensemble ML and the necessity for mechanistic insight, this approach offers a clearer mapping of the combinatorial genetic processes underlying complex diseases. While applied here to AAV, the method is highly adaptable for exploring the genetic architecture of diverse populations and complex traits.

19
A Multi-Agent Large Language Model Reasoning Engine for Early Detection of Pediatric Growth Disorders

Rabbani, N.; Mettner, J.; Lee, K.; Soto-Rivera, C. L.; Windberger, A.; Santiago, K.; Hatoun, J.; Correa, E. T.; Vernacchio, L.; Kohane, I.

2026-08-31 health informatics 10.64898/2026.08.28.26361655 medRxiv
Top 0.6%
1.7%
Show abstract

Routine childhood growth surveillance is a cornerstone of pediatric care. Growth pattern abnormalities are often early manifestations of chronic disease. Yet subtle abnormalities are frequently underrecognized, leading to diagnostic delays and avoidable morbidity. We introduce SPROUT (System for Pediatric Recognition Of Undiagnosed Trajectories), a generalized, multi-agent large language model (LLM) reasoning system designed to identify a broad spectrum of pediatric growth-related conditions from longitudinal electronic health records (EHRs) earlier than standard clinical practice. Using a large pediatric primary care EHR dataset, we developed and validated SPROUT as a two-stage system. First, a highly specific LLM screener flags concerning longitudinal growth patterns. Second, an Orchestrator module coordinates a multidisciplinary panel of LLM agents to generate a ranked differential diagnosis. To correct systemic reasoning errors, a Trainer module injects meta-knowledge into the panel via a dedicated "Learner" agent. Diagnostic capability was evaluated using a walk-forward, visit-by-visit simulation leading up to the diagnosis date. The SPROUT screener model achieved 98% (83/85) specificity and 28% (9/32) sensitivity on a gold-standard dataset of pediatric primary care patients when evaluated one year before the index date, and 100% specificity and 47% sensitivity when evaluated using longitudinal data up to the day of diagnosis. When applied to 300 control patients (i.e., healthy or undiagnosed), the screener flagged 15. Subsequent expert panel review confirmed high suspicion for undiagnosed pathology in 33% (5/15) of these cases. In chronological walk-forward validation on disease cases, the diagnostic engine identified conditions well before standard-of-care documentation. One year prior to clinical diagnosis, the system achieved sensitivities of 81% for type 1 diabetes mellitus, 56% for pituitary disorders, and 44% for celiac disease. The SPROUT multi-agent system demonstrates the ability to detect a significant portion of latent growth-related pediatric conditions months to years before current clinical standards while minimizing false positives. These results support its potential as a decision support tool for reducing diagnostic delays in pediatric care.

20
Age at Onset and Liability to Disorder: Estimating Covariances in Censored Populations

Neale, M. C.; Maes, H. H.; Mullins, L. K.; Singh, M.; Balbona, J.; Kirkpatrick, R. M.; Brick, T. R.; Hunter, M. D.; Boker, S. M.; Castro-de-Araujo, L.; Schork, A. J.; Krebs, M. D.; Mefford, J. A.

2026-08-06 genetic and genomic medicine 10.64898/2026.08.04.26359043 medRxiv
Top 0.7%
1.5%
Show abstract

Studies of resemblance for disorders and other traits measured at the binary (yes/no) level between relatives frequently contain individuals who are currently in the negative category but who will become positive in future. For example, a 10-year-old may develop depression in the future, but is as yet unaffected. Such censoring can substantially bias estimates of correlation between relatives. To overcome this problem we develop a model for the association between liability to a disorder, and its age at onset. The model is designed for data from pairs of relatives to enable estimation of the correlation between an individuals' liability to disorder and their age at onset. Usually, such information is not available at the individual level, because age at onset is uniquely available when onset has occurred. Lacking variation in disorder status, data from non-related persons cannot estimate the covariance between liability and age at onset. Data from relatives can resolve this issue when there is a correlation in liability between the relatives, because different age at onset distributions would be expected in concordant vs. discordant pairs of relatives. Greater severity and worse outcomes are often observed among those with earlier onset, so a correlation between disorder liability and age at onset seems likely in many cases. In this article we present the basic theory of the model, implemented as a mixture distribution, and an application to cannabis use in a Virginia Twin Study of Adolescent Behavioral Development. A negative association of (-.212) between age at onset an liability was found, with confidence intervals of -.263 to -.152, which do not cross zero. The method contrasts with Cox Proportional Hazards, in which disorder liability and onset timing are treated as a single dimension.